Speech Recognition Using Energy Parameters to Classify Syllables in the Spanish Language
نویسندگان
چکیده
This paper presents an approach for the automatic speech recognition using syllabic units. Its segmentation is based on using the ShortTerm Total Energy Function (STTEF) and the Energy Function of the High Frequency (ERO parameter) higher than 3,5 KHz of the speech signal. Training for the classification of the syllables is based on ten related Spanish language rules for syllable splitting. Recognition is based on a Continuous Density Hidden Markov Models and the bigram model language. The approach was tested using two voice corpus of natural speech, one constructed for researching in our laboratory (experimental) and the other one, the corpus Latino40 commonly used in speech researches. The use of ERO parameter increases speech recognition by 5% when compared with recognition using STTEF in discontinuous speech and improved more than 1.5% in continuous speech with three states. When the number of states is incremented to five, the recognition rate is improved proportionally to 97.5% for the discontinuous speech and to 80.5% for the continuous one.
منابع مشابه
Algorithms and Methods for the Automatic Speech Recognition in Spanish Language using Syllables
This work examines the results of incorporating into Automatic Speech Recognition the syllable units for the Spanish language. Because of the boundaries between phonemes-like units its often difficult to elicit them; the use of these has not reached a good performance in Automatic Speech Recognition. In the course of the developing the experiments three approaches for the segmentation task were...
متن کاملSpoken Term Detection for Persian News of Islamic Republic of Iran Broadcasting
Islamic Republic of Iran Broadcasting (IRIB) as one of the biggest broadcasting organizations, produces thousands of hours of media content daily. Accordingly, the IRIBchr('39')s archive is one of the richest archives in Iran containing a huge amount of multimedia data. Monitoring this massive volume of data, and brows and retrieval of this archive is one of the key issues for this broadcasting...
متن کاملA Pragmatic Study of Speech Acts by Iranian and Spanish Nonnative English Learners
This study was an attempt to investigate Iranian and Spanish intermediate nonnative English learners’ request strategies to their faculty. To this aim, 74 (50 Iranian and 24 Spanish) nonnative English intermediate learners participated in this study. A discourse completion test (DCT) was used to elicit the request strategies used by the participants. The findings suggested the participants empl...
متن کاملExtraction and representation of prosodic features for language and speaker recognition
In this paper, we propose a new approach for extracting and representing prosodic features directly from the speech signal. We hypothesize that prosody is linked to linguistic units such as syllables, and it is manifested in terms of changes in measurable parameters such as fundamental frequency (F 0), duration and energy. In this work, syllable-like unit is chosen as the basic unit for represe...
متن کاملClassification of emotional speech using spectral pattern features
Speech Emotion Recognition (SER) is a new and challenging research area with a wide range of applications in man-machine interactions. The aim of a SER system is to recognize human emotion by analyzing the acoustics of speech sound. In this study, we propose Spectral Pattern features (SPs) and Harmonic Energy features (HEs) for emotion recognition. These features extracted from the spectrogram ...
متن کامل